Papers with holistic evaluation
Gentopia.AI: A Collaborative Platform for Tool-Augmented LLMs (2023.emnlp-demo)
Copied to clipboard
Binfeng Xu, Xukun Liu, Hua Shen, Zeyu Han, Yuhan Li, Murong Yue, Zhiyuan Peng, Yuchen Liu, Ziyu Yao, Dongkuan Xu
| Challenge: | Existing frameworks for Augmented Language Models lack flexibility, democratization, and holistic evaluation. |
| Approach: | They propose a lightweight and extensible framework for Augmented Language Models called Gentopia. |
| Outcome: | The proposed framework integrates language models, task formats, prompting modules, and plugins into a unified paradigm. |
Process Evaluation for Agentic Systems (2026.findings-eacl)
Copied to clipboard
| Challenge: | Recent adoption of LLM-based assistants has led to premature assumptions about their reliability and general capability. |
| Approach: | They propose to assess the feasibility of automatic process evaluation for critical applications such as medicine, finance, law and infrastructure. |
| Outcome: | The proposed evaluations are based on a small-scale study to assess the feasibility of automated process evaluation, present a compliance score, analyse use cases of bad and good behaviours, and offer recommendations for more holistic evaluation. |